Data Science
A US Census Report on Noncitizen Voting Used Bad Data to Reach Faulty Conclusions
Trump has touted a recent Census report. WIRED found grave flaws in its analysis and the process behind it, and confirmed the identity of several of its authors--among them a one-time DOGE affiliate. A report published by the US Census Bureau last month purported to uncover evidence backing up President Donald Trump's baseless claims that noncitizens voting cost him the 2020 election . It was the latest salvo in Trump's broader assault on the safety of US elections. "I WON THE ELECTION," Trump quickly declared on Truth Social after the publication of the report.
AIhub monthly digest: July 2026 โ time-series anomaly detection, music generation, and RoboCup in action
Welcome to our monthly digest, where you can catch up with any AIhub stories you may have missed, peruse the latest news, recap recent events, and more. This month, we find out about time-series anomaly detection, delve into music generation, honour award winners, and catch up on the action from the RoboCup humanoid soccer league. We caught up with Thi Kieu Khanh Ho to find out more about her work on time-series anomaly detection, what inspired her to study AI, and what she plans to work on next. This interview is part of our series featuring the AAAI Doctoral Consortium participants. In the latest in our series of IJCAI interviews, AIhub ambassador Liliane-Caroline Demers spoke to Franรงois Pachet to find out more about his work on music generation with AI.
Interview with Thi Kieu Khanh Ho: Time-series anomaly detection
The latest interview in our series with the AAAI/SIGAI Doctoral Consortium participants features Thi Kieu Khanh Ho who is studying time-series anomaly detection. We found out more about her research, and what inspired her to study AI, and what she plans to work on next. Tell us a bit about your PhD -- where are you studying, and what is the topic of your research? I am doing my PhD at McGill University and Mila - Quรฉbec AI Institute, in the Department of Electrical and Computer Engineering, supervised by Professor Narges Armanfard. My research focuses on time-series anomaly detection, the problem of teaching AI systems to recognize when something unusual or abnormal is happening in complex, real-world data streams, without relying on large amounts of labeled examples.
Learning Effective Soliton Dynamics from Scattering Data
Minor, Seth, Dukic, Vanja, Bortz, David M.
In such settings, the inverse scattering transform (IST) of Ablowitz, Kaup, Newell, and Segur [2] has enjoyed a rich and successful history, and is now the standard theoretical framework for deriving reduced-order evolution equations for soliton dynamics. Although these derivations are traditionally of an analytical - rather than data-driven - nature, recent work has employed the IST formalism as a tool for experimental data analysis, using the technique to analyze soliton content from empirical measurements [8, 15, 24]. Moreover, recent approaches using alternative parameterization techniques have demonstrated that the learning of reduced-order, interpretable equations of motion for solitons is tenable in a data-driven setting [6, 26, 27]. Despite the success of this recent work, however, little effort has been devoted to developing a data-driven modeling approach based on the IST itself, most likely due to the fact that the framework is fundamentally problem-specific. In this paper, we address the question of whether effective soliton dynamics can be inferred directly from observed scattering data (as opposed to being derived or approximated analytically).
Cloudflare will filter out web crawlers that serve AI companies
The hosting platform wants sites to have more control over how AI companies use their content. Cloudflare has announced plans to automatically block mixed-use web crawlers that index websites for search engines and act as AI agents and trainers at the same time. The company previously offered its customers the optional ability to prevent crawlers from scraping their sites for AI chatbots, but now Cloudflare's stance is becoming more defensive by default. Now that the majority of traffic on the Internet is non-human, we must go further and act faster so that a sustainable ecosystem can emerge, Matthew Prince, Cloudflare's CEO and co-founder shared in a statement. Cloudflare's new tools and partnerships give website owners increased visibility and commercial opportunities and benefit AI companies that have bots with clear and transparent intent.
Approximate full-conformal multi-task regression with reproducing kernels
Razafindrakoto, Davidson Lova, Celisse, Alain, Lacaille, Jรฉrรดme
Multi-task regression aims at jointly solving multiple regression problems, called tasks. Compared to solving each task separately, better performances can be achieved as long as the tasks are sufficiently related. Full-conformal prediction is a framework that formulates a data-dependent prediction-region containing the unknown output-vector at any prescribed confidence level. However, explicit computation of this prediction-region is intractable in general since it requires training infinitely many predictors. The present work focuses on multi-task regression in a Reproducing Kernel Hilbert Space (RKHS) of vector-valued functions. This computational issue is addressed by designing an approximating predictionregion containing the full-conformal one. This construction is carried out in two scenarios: piq when the inter-task covariance-matrix is known, and piiq when this matrix is estimated. In terms of volume, the tightness of this approximation is assessed theoretically by means of an upper-bound in the first scenario. It is also empirically proved to improve upon the split-conformal prediction on synthetic data in both scenarios.
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions
Li, Mingyi, Tsuchiya, Taira, Yamanishi, Kenji
We study policy optimization for online episodic tabular Markov decision processes with unknown transition kernels, aiming for best-of-both-worlds guarantees together with data-dependent regret bounds. Recent work (Dann et al., 2023; Li et al., 2026) has shown that policy optimization can adapt to both adversarial and stochastic losses with first-order, second-order, and path-length bounds, but only under known transitions, leaving open whether such data-dependent guarantees are achievable by policy optimization when the transition kernel is unknown. We resolve this by developing a new algorithm based on optimistic follow-the-regularized-leader that attains these guarantees under unknown transitions. The key ingredient is a new design of optimistic $Q$-function estimators together with a data-dependent transition bonus that controls estimator bias through the loss-prediction error. Our analysis further identifies an unavoidable transition-dependent complexity term that captures the intrinsic cost of estimating the transition kernel. As a result, we obtain first-order, second-order, and path-length bounds with the transition-dependent complexity term while simultaneously achieving gap-dependent $\mathrm{polylog}(T)$ regret in the stochastic regime.
Connectivity Estimation using Stochastic Graph Heat Modelling
Goerttler, Stephan, Wu, Min, He, Fei
A growing number of techniques leverage the spatial structures that underlie many real-world datasets. Despite these advances, the complementary task of estimating spatial structures and understanding their role within these techniques has often been overlooked. In neurophysiological data analysis specifically, numerous methods exist to estimate brain connectivity, but most are not explicitly model-based, dynamic, multivariate, or directed. To address these limitations, we previously introduced noise-driven heat modelling on graphs for neurophysiological connectivity estimation. In this study, we extend this framework by relaxing earlier noise assumptions and adding regularisation to improve robustness. We also develop a simulation procedure to characterise and evaluate our technique in a controlled setting. Finally, we demonstrate that the technique is able to capture meaningful spatial structure across two experiments, each using two real-world datasets. The explicit model formulation of our connectivity estimator has the potential to improve the interpretability of graph-based techniques across a wide range of applications. The code implementing our method is available at https://github.com/sgoerttler/Heat_Connectivity.
What Drives the Inlier-Memorization Effect? A Theory of Outlier Detection via Early Training Dynamics
Outlier detection (OD) aims to identify anomalous instances by learning the underlying structure of normal data (inliers), and is particularly challenging in fully unsupervised settings where no information about anomalies is available during training. Recent advances have leveraged the inlier-memorization (IM) effect, a phenomenon in which deep models memorize inlier patterns earlier than those of outliers, as a powerful signal for distinguishing outliers. However, despite its empirical success, the theoretical understanding of the IM effect remains limited. In this work, we present a theoretical study of the IM effect. Focusing on a simple autoencoder, we show that, under mild assumptions, the model can successfully memorize inliers while failing to memorize outliers during certain stages of early training. In particular, we characterize not only the emergence of the IM effect, but also its strength and persistence, and analyze how these properties depend on the data distribution and parameter initialization. In addition, building on these insights, we derive simple yet practical guidelines for enhancing the IM effect, including data preprocessing and parameter initialization schemes, achieving state-of-the-art performance on the ADBench datasets. Our findings provide a theoretical foundation for the IM effect and offer actionable directions for improving IM-based outlier detection methods.